Proteins: Structure, Function, and Bioinformatics
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match Proteins: Structure, Function, and Bioinformatics's content profile, based on 88 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Ghosh, B.; MUKHERJEE, A.
Show abstract
Short peptides pose distinct challenges for computational structural biology due to their lack of stable tertiary structures, high conformational flexibility, and limited evolutionary signals. To address how modern deep-learning architectures navigate these challenges, we conducted a comprehensive benchmarking of five state-of-the-art protein structure prediction models: AlphaFold2, RoseTTAFold2, ESMFold, OmegaFold, and DMPfold2. Using a curated dataset of experimentally determined short peptide structures (10-49 amino acids) from the Protein Data Bank, we systematically evaluated predictive performance across varying sequence lengths and secondary structure classes. Our results demonstrate that prediction accuracy systematically improves with peptide length. Furthermore, all models perform significantly better on -helical and mixed-structure peptides compared to {beta}-sheet-rich and intrinsically disordered sequences. Among the evaluated methods, AlphaFold2 and the single-sequence language models, ESMFold and Omegafold proved to be the most consistent and accurate overall. We also observed that internal model confidence scores are imperfectly calibrated for short peptides, necessitating cautious interpretation. Finally, by extending our analysis to the dbAMP3 dataset of uncharacterized antimicrobial peptides, we demonstrate that a multi-model consensus approach provides a rational framework for identifying robust structural hypotheses in the absence of experimental reference structures.
Zhang, S.; Wang, Z.; Chen, E.
Show abstract
Human mitochondrial ATP synthase is an essential rotary motor enzyme that produces most of the cellular ATP through oxidative phosphorylation. Its membrane-embedded Fo sector contains highly hydrophobic transmembrane subunits that are challenging to study in aqueous environments without detergents. This study explores whether applying the QTY code can reduce the hydrophobicity of selected ATP synthase Fo subunits while preserving their overall molecular structures. We applied the QTY code to eight human ATP synthase Fo subunits: ATP6, ATP8, ATPK, ATP68, ATPMK, AT5G1, AT5G2, and AT5G3. Hydrophobic amino acids leucine (L), isoleucine (I), valine (V), and phenylalanine (F) in transmembrane regions were systematically replaced with hydrophilic glutamine (Q), threonine (T), and tyrosine (Y). Four native subunits with available CryoEM structures from human ATP synthase (PDB: 8H9S) were superposed with their AlphaFold3-predicted QTY analogs. The native ATP synthase Fo subunits superposed well with their respective QTY analogs. For the CryoEM-native comparisons, RMSD values ranged from 0.565[A] to 2.546[A]. For the AlphaFold3-native comparisons of subunits without CryoEM structures, RMSD values ranged from 0.204[A] to 0.297[A]. Despite substantial QTY substitutions in the transmembrane regions, ranging from 38.89% to 50.79%, the QTY analogs retained similar overall folds, molecular weights, and isoelectric points. Hydrophobic surface analysis showed that the QTY analogs had reduced hydrophobic patches compared with their native counterparts, with average hydrophobicity decreasing from 0.2959 in native proteins to -1.1023 in QTY analogs. These structural bioinformatics studies suggest that the QTY code can be applied to ATP synthase Fo subunits to generate more hydrophilic, potentially water-soluble analogs while preserving overall structural similarity. These results extend the application of the QTY code to the membrane-embedded Fo sector of ATP synthase and provide a foundation for future experimental studies testing whether these QTY analogs can be expressed, purified, and evaluated for assembly or proton-transfer-related functions.
Marien, J.; Sritharan, S.; Caviglia, B.; Versini, R.; Auclair, L.; Barraud, P.; Tisne, C.; Reguei, A.; Murail, S.; Leclerc, F.; Duboue-Dijon, E.; Tao, J.; Basdevant, N.; Baaden, M.; Prevost, C.; Sacquin-Mora, S.; Taly, A.
Show abstract
The advent of deep learning-driven tools such as AlphaFold has revolutionized the prediction of biomolecular structures, offering unprecedented accuracy and accessibility for proteins, RNA, and their complexes. While these tools have demonstrated remarkable success in benchmarking competitions and enabled experimentalists to generate models with ease, their widespread use has also highlighted persistent challenges. These include difficulties in assessing model confidence, limitations in predicting transmembrane domains, nucleic acids, conformational diversity, and interactions with ions or ligands, as well as the tendency to misfold intrinsically disordered regions (IDRs). In this perspective, we critically evaluate the strengths and limitations of current AI-based structure prediction tools through illustrative examples with a particular emphasis on the impact of explicit ion modelling. We notably report how the explicit addition of a few potassium cations to the prediction of IDRs or G-quadruplexes can trigger massive conformational switches compared to "dry" predictions. On this basis, we suggest modelling sequences both "dry" and in the presence of explicit potassium cations as a simple, practical way to sample alternative conformations and to expose disordered regions that current predictors tend to over-fold. We discuss the importance of reporting confidence metrics in publications to avoid overinterpretation. Furthermore, we address the unique challenges of RNA structure prediction, where data scarcity and structural complexity limit the performance of both classical and deep learning methods. Our analysis underscores the need for continued methodological advancements, integration of complementary computational tools, and expansion of high-quality experimental datasets. TOC Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/740587v1_ufig1.gif" ALT="Figure 1"> View larger version (16K): org.highwire.dtl.DTLVardef@fc2f42org.highwire.dtl.DTLVardef@82c1a4org.highwire.dtl.DTLVardef@7733cborg.highwire.dtl.DTLVardef@1e97db2_HPS_FORMAT_FIGEXP M_FIG C_FIG
Ito, F.; Konishi, M.; Nakamura, R.; Akizawa, T.
Show abstract
The development of small synthetic catalytic peptides, or "catalytides," offers a promising therapeutic strategy for the targeted degradation of amyloid-beta (A{beta}). Among these, the pentapeptide SKGQA mimics the proteolytic activity of serine proteases despite its minimal size. However, the molecular mechanism enabling such a short peptide to achieve effective cleavage at multiple sites remains unclear. In this study, we utilized HADDOCK docking and molecular dynamics (MD) simulations to investigate the interaction between SKGQA and the A{beta}(17-42) region. Our results demonstrate that SKGQA operates through a highly dynamic process, where the substrate serves as a scaffold to stabilize "serine protease-like" active geometries from a flexible conformational ensemble. We identified distinct "stable binding" and "stochastic attack" modes, explaining the peptides ability to facilitate both high-probability and multi-site cleavage. Given its minimal size, SKGQA may also benefit from enhanced accessibility to dense amyloid environments compared to larger proteases. These findings provide a fundamental understanding of minimal enzymatic function and offer a transformative platform for designing next-generation, cost-effective catalytides.
Jain, S.; Mehta, N. K.; Raina, S.; Kumar, P.; Varun, ; Raghava, G. P. S.
Show abstract
While most existing methods are limited to predicting the tertiary structures of proteins containing only canonical residues, the PEPstrMOD server (developed in 2015) pioneered structure prediction for chemically modified and non-natural peptides. Despite its widespread use, the original framework was restricted to peptides of 7 to 25 residues and relied on older backbone-prediction algorithms. To address these limitations, we present PEPstrMOD2, which introduces three major advancements over its predecessor. First, it replaces the original in-house coordinate generation with state-of-the-art deep learning (DL) algorithms, leveraging AlphaFold2 and ESMFold for highly accurate initial structure prediction. Secondly, it greatly expands the accessible chemical space through incorporation of new, AMBER force-field compatible library of 257 post-translational modifications (PTMs), 428 non-canonical amino acids (NCAAs), and 243 terminal modifications. Lastly, through the application of native scalability of AlphaFold2 (AF2) and ESMFold (EF), PEPstrMOD2 eliminates the original restrictions of the length, enabling the structural modeling of longer, complex therapeutic peptides and small proteins. We evaluated the performance of PEPstrMOD2 against state-of-the-art methods across three distinct peptide datasets. For the AfCyc dataset consisting of 80 cyclic peptides, PEPstrMOD2 obtained a competitive average atom-level Root Mean Square Deviation (RMSD) of 2.05 angstroms, compared to 1.13 angstroms by AlphaFold3 (AF3) and 1.82 angstroms by AfCycDesign. Remarkably, for the modified peptide ModPep433 dataset, PEPstrMOD2 outperformed AF3, achieving the lower average RMSD score of 4.49 angstroms against 4.67 angstroms of AF3. Furthermore, in the case of the ModPep16 benchmark, PEPstrMOD2 achieved 2.50 angstroms average RMSD value, which is two times more accurate than that of the original PEPstrMOD (5.84 angstroms). In summary, PEPstrMOD2 provides a powerful, high-throughput, and highly accurate platform to facilitate peptide-based drug development and structural biology research. While the original PEPstrMOD was restricted to a web server interface, PEPstrMOD2 is available as both an intuitive webserver and a standalone command-line tool via GitHub, featuring Docker support for easy deployment and reproducible, large-scale modeling pipelines (https://webs.iiitd.edu.in/raghava/pepstrmod/).
Zhang, S.; Xiao, E.
Show abstract
Human aquaporins (AQPs) are essential membrane channels, yet their inherent hydrophobicity complicates structural and functional studies. We present the systematic application of the QTY code to human AQPs, integrating it with AlphaFold 3 structure prediction to design and validate that four-representative human AQPs (AQP1, AQP3, AQP4, AQP7) can be converted into water-soluble analogs while maintaining their conformation. This approach features a novel platform for editing challenging membrane proteins. The QTY code was applied to the transmembrane regions of the selected four AQPs. Subsequently, the water-soluble QTY analogs of the four AQPs were predicted using AlphaFold 3. The predicted structures were superposed with CyroEM- or X-ray-determined native structures in PyMOL. Further analyses included root-mean-square deviation (RMSD) calculations, visualization of hydrophobic surface reduction, and inspection of conserved protein-ligand binding ability. After applying the QTY code, sequence changes between native AQPs and their QTY analogs was significant (42.86-48.80%). Nevertheless, their structures superposed well in analyses, with only slight deviations (RMSD < 0.6 [A]). In addition, the surface hydrophobicity of all QTY-edited AQPs was significantly reduced. Importantly, molecular contacts between the cholesterol ligand and protein were largely preserved for both native AQP1 and its QTY analog. Finally, all AlphaFold3-predicted structures for AQPs have high confidence values (pLDDT > 90; pTM ~0.83), supporting the reliability of the predicted structures. The findings demonstrate that membrane protein hydrophobicity can be edited and reduced without compromising fold integrity or functional architecture. Integration of the QTY code with AlphaFold 3 affords a high-throughput platform for designing water-soluble, structurally faithful analogs of challenging membrane proteins. Such a strategy can provide a potent platform for detergent-free biochemical studies and water-soluble analogs for therapeutic monoclonal antibody discoveries, thus advancing research of this pharmacologically important protein family.
Yadav, L. R.; Chauhan, S. B.; Joshi, M.; Mande, S. C.
Show abstract
Ribonucleotide reductases (RNRs) employ radical chemistry to generate deoxyribonucleotides required for DNA synthesis and repair. A notable feature of RNRs is half-site reactivity, where, despite the enzyme being a symmetric 2 dimer, only one active site is catalytically active at a time while the other remains in a "poised" state for substrate binding. This phenomenon is tightly linked to the asymmetric 2{beta}2 interaction required for radical transfer. Here, we determined cryo-EM structures of the -subunit in the apo and holo states, i.e., the complex bound to TTP (effector) and GDP (substrate). The structures reveal asymmetric binding of the effector TTP and the substrate GDP across the dimer, with concomitant stabilization of loops surrounding the ligand-binding site. Interestingly, this asymmetry leads to well-resolved N-terminal density for [~]150 residues in the substrate-bound subunit, but weak density for this region in the effector-bound monomer. N-terminal domains are unresolved in both monomers of the apo structure. Isothermal titration calorimetry supports asymmetric binding of pyrimidine effectors with micromolar affinities. Molecular dynamics simulations and three-dimensional variability analysis reveal synchronous motions of loop 2, which together with the N-terminal domain drive alternate opening and closing of the active sites in the two monomers. These conformational dynamics provide key insights into the mechanistic basis of half-site reactivity. Together, these findings provide new insights into the structural dynamics and thermodynamic principles governing regulation and half-site activity in Class Ib RNRs. Significance statementRibonucleotide reductases (RNRs) are essential enzymes that supply the building blocks required for DNA synthesis and repair, yet the structural basis of their half-site reactivity has remained unclear. Using cryo-electron microscopy, calorimetry, molecular dynamics simulations, and conformational variability analysis, we show that the catalytic -subunit of a Class Ib RNR exhibits asymmetric nucleotide binding and coordinated conformational dynamics between the two monomers. These motions drive alternating opening and closing of the active sites and are linked to differential stabilization of the N-terminal region. Our findings suggest that asymmetric conformational gating and N-terminal sampling regulate productive interaction with the radical-generating {beta}-subunit, providing a mechanistic framework for understanding half-site reactivity and allosteric regulation in RNRs.
Orr, A. K.; Bateman, A.
Show abstract
MotivationSpurious protein sequences, resulting from gene prediction errors, theoretically should not yield folded structures. AlphaFold2 was previously shown to predict short spurious sequences with high pLDDT scores and was therefore unlikely to distinguish between real proteins and spurious proteins which are usually short. We evaluate whether newer structure prediction methods (ESMFold and AlphaFold3) similarly predict short sequences with high pLDDT or if they better discriminate between spurious and real proteins. ResultsAll three structure prediction methods (ESMFold, AlphaFold2, and AlphaFold3) predict short spurious sequences from AntiFam with unexpectedly high pLDDT scores, however the discrimination between spurious and real proteins improves beyond 100 amino acids. By analysing sequences with disparate pTM and pLDDT scores, we identified two likely spurious shadow ORFs in Swiss-Prot and one potentially non-spurious AntiFam entry. Using the structure prediction scores, we developed a Gaussian Process Model and evaluated its performance on AlphaFold DB, identifying potential spurious proteins at scale. While limited on its own, this model can increase confidence in spurious protein identification when combined with other methods. AvailabilityStructure predictions are available at https://doi.org/10.5281/zenodo.18390113. Model implementation and figure generation code are available at https://github.com/0rra/fold_unfold2.
Dohmen, R. L.; Hoogerwerf, G.; Xie, A.; Hoff, W. D.
Show abstract
A universal mechanism in molecular evolution is functional and structural divergence of members of a protein family. The ability of AlphaFold to predict atomic-resolution protein structures promises to accelerate insights into this process. We study the interplay of changes in sequence, structure, and function in photoactive yellow protein (PYP), a family of bacterial blue light photoreceptors. Halorhodospira halophila contains two PYP homologs that diverged to 60% sequence identity, differ 100-fold in the lifetime ({tau}pB) of their pB signaling intermediate, and display altered peak wavelengths ({lambda}max) for color sensing. We resurrected ancestral PYPs and determined these properties along the resulting recapitulating evolutionary divergence. The resurrected ancestral PYP is functionally similar to PYP1, indicating divergence on the path to PYP2. AlphaFold predictions for PYP2 and these ancestral proteins revealed the absence of structural changes compared to the crystal structure of PYP1. To experimentally validate these predictions, we optimized second-derivative Fourier transform infrared (FTIR) spectroscopy. The FTIR spectra of PYP1 and 2 and their resurrected ancestral proteins demonstrated clear differences in their secondary structure. These results demonstrate an important limitation of AlphaFold and show how ancestral sequence reconstruction combined with spectroscopic approaches yields insights into divergence in a protein family.
Stephenson, H.; Voicu, D.; Novakov, V.; Levy, M.; Marsilio, J.
Show abstract
With the growing use of machine-learning-assisted pipelines for designing, characterizing, and optimizing biomolecules, the reliability of structure prediction models is increasingly important. PolyFold is a benchmarking framework developed to evaluate open-use structure prediction models, Boltz-2 and OpenFold 3, as commercially accessible alternatives to AlphaFold 3. We outline an end-to-end workflow automation tool to streamline input file creation, batch automation, and comprehensive analysis of model outputs for leading open-use structure prediction models. We curated an evaluation dataset of several thousand high-quality Protein Data Bank structures, homology-filtering against the training sets of both models to ensure a fair analysis. We then implemented an evaluation pipeline incorporating structural metrics (RMSD, TM-score, lDDT, etc.), interface metrics (DockQ, ilDDT, iRMSD, etc.), and physicochemical realism checks (based on bond lengths, angles, molecular internal energies, etc.). We identify key performance disparities, observing that Boltz-2 is generally superior to OpenFold 3, though the differential is partially attributable to residual homology leakage not accounted for by prevailing test set curation practices. We thus recommend a new method for homology-reducing when building a test set using length-weighted average fractional identity cutoffs rather than lowest chain fractional identity cutoffs. Even in eliminating residual leakage, Boltz-2 still performs better on full-set comparisons and a variety of important partitions (nucleic acids, protein-ligands, Ab-Ags, etc.). Both models are strong at folding monomeric structures, though struggle with homomultimer placement and small molecule physical realism, demonstrating enduring limitations of machine learning methods. This work is the first end-to-end, open-use, and reproducible platform for systematically assessing state-of-the-art structure prediction models. PolyFold enables practitioners to determine how models compare in performance on specific inference tasks and supports the broader adoption of accessible computational tools to facilitate biomolecular science.
Bailey, J. S.; Phan, N.; Spina, S. C.; Getman, R. B.; Paulson, J. A.; Kimmel, B. R.
Show abstract
Protein-RNA complexes drive fundamental cellular processes such as transcription and translation. Despite the prevalence and importance of protein-RNA interactions, the field lacks reliable and accessible methods to quantify the energetic favorability of these interactions. We propose an experimentally tuned protein-RNA score function that can be directly implemented into ROSETTA. Fine-tuning these score functions for predictive tasks requires repeated evaluations on a set of protein-RNA complexes, which can be computationally expensive given the number of parameters to tune. We used Bayesian Optimization to efficiently improve the energetic agreement between ROSETTA and experimentation. We observe significant interactions for specific RNA subclasses, serving as further confirmation of the physical validity of the score function. Beyond protein-RNA interaction prediction, we establish a framework to efficiently fine-tune ROSETTA score functions for any protein-class interaction using Bayesian Optimization. TOC FIGURE O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=100 SRC="FIGDIR/small/739244v1_ufig1.gif" ALT="Figure 1"> View larger version (18K): org.highwire.dtl.DTLVardef@700559org.highwire.dtl.DTLVardef@6f4ca6org.highwire.dtl.DTLVardef@11157c0org.highwire.dtl.DTLVardef@1981456_HPS_FORMAT_FIGEXP M_FIG C_FIG
Ye, M.; Wang, Y.-H.; Brogi, M.; Parks, J. M.; Kuo, K. M.; Gumbart, J. C.
Show abstract
Protein structure predictors achieve high single-state accuracy, but it remains unclear whether they can recover functionally relevant conformational ensembles or account for the presence of ligands and/or binding partners. Here, we benchmark AlphaFold3, Boltz-2, Chai-1, and BioEmu on four canonical multi-state proteins (Pf-MATE, LAO, SecA, and {beta}2AR), quantifying state bias and sampling breadth against experimental reference structures. Models frequently default to a dominant state represented in the PDB; small-molecule ligands have weak or inconsistent effects, while large protein partners drive clear conformational switching between states. Multiple sequence alignment (MSA)-based approaches (AF-Cluster and random subsampling) recapitulate similar biases, indicating that this behavior is not unique to newer architectures. These results underscore current limitations for multi-state protein structure prediction and structure-guided ligand discovery. TOC Graphic O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=111 SRC="FIGDIR/small/737860v1_ufig1.gif" ALT="Figure 1"> View larger version (12K): org.highwire.dtl.DTLVardef@3bf389org.highwire.dtl.DTLVardef@1f1c436org.highwire.dtl.DTLVardef@188ea8aorg.highwire.dtl.DTLVardef@1de236e_HPS_FORMAT_FIGEXP M_FIG C_FIG
Kolypetris, G.; Djurabekova, A.; Lasham, J.; Simsive, L.; Vonck, J.; Sharma, V.
Show abstract
Cryogenic-electron microscopy (cryo-EM) has revolutionized the field of protein structural biology. The structures of large membrane proteins are now routinely determined by cryo-EM to near atomic resolution. However, in the medium resolution range of cryo-EM maps (>[~]2 [A]), negatively charged sidechains of acidic residues are not well-resolved due to the negative electrostatic potential of the region. This may lead to incorrect sidechain models for residues like glutamic acid or aspartic acid that are central for proton transfer activity in various respiratory and photosynthetic enzymes. We previously proposed that the acidic residues with weak or non-existent cryo-EM density can be modeled to represent their low proton affinity conformations. Here, we tested this hypothesis on a larger data set of acidic amino acid residues in two high-resolution respiratory complex I structures. By using faster sidechain modeling and proton affinity prediction tools, we created a workflow that generates sidechain conformations of selected amino acid residues. We validated the sidechain conformation predictions by Q-score analysis and atomistic molecular dynamics simulations in different charged states. The proposed workflow provides a way to rapidly obtain sidechain conformations of acidic residues with weak cryo-EM densities and can be integrated into the existing cryo-EM modeling pipelines to speed up sidechain rotamer prediction.
Senguler Ciftci, F.; Erman, B.
Show abstract
Quantifying how cooperative, many-body relationships drive allostery in protein networks remains a major challenge. To address this, we develop the Laplacian minor hierarchy, a mathematical framework that characterizes the geometric invariants of a protein network. Lower-order minors yield standard metrics including the partition function and effective distances, whereas higher-order minors define novel topological measures: cooperation indices, each bounded between zero and one, that characterize pathway correlations at increasing levels of complexity, the third-order minor determines whether allosteric pathways are correlated or uncorrelated, and the fourth-order minor quantifies how distinct pathways communicate through intermediary residues. We apply this framework to analyze the evolutionary adaptation of the PSD95pdz3 domain from Class I to Class II ligand specificity via mutations G330T and H372A. The cooperation index demonstrates a distinct evolutionary hierarchy: the G330T mutation establishes distributed pathway couplings that the H372A mutation subsequently exploits, whereas H372A alone produces minimal global changes. Furthermore, the fourth-order analysis identifies His317 as a critical intermediary node bridging the class-switching (330-372) and class-bridging (330-400) allosteric pathways. These results demonstrate that allosteric dependencies emerge only when mutations accumulate in specific combinations, with a hierarchical organization of pathways structured around position 330 and intermediary nodes His317 and Phe400. Rather than predicting allosteric mechanisms, this framework provides a mechanistic explanation for why and how allostery emerges during protein evolution.
Jindal, M.; Mahato, R.; Das, S.; Guha, A.; Majila, K.; Arvindekar, S.; Vaidya, A. T.; Viswanath, S.
Show abstract
The Mitochondrial contact site and Cristae Organizing System (MICOS) complex is an inner mitochondrial membrane (IMM) assembly present at the cristae junction. It is responsible for regulating cristae formation and remodeling. However, its structure is not known. We applied Bayesian integrative structure determination to characterize the structure of the Mic60, Mic19, Mic10, and Mic13-containing MICOS complex combining AlphaFold predictions with data from crosslinking mass spectrometry, biochemical assays, electron tomography, homology modeling, and sequence alignments. The integrative structure revealed novel mutual interfaces among Mic10N,C, Mic60LBS1,LBS2,mitofilin, and Mic13central,C, which were experimentally validated. Several likely-pathogenic missense mutations also localize to these novel interfaces, highlighting their importance. Our results indicate that Mic13 likely facilitates MICOS assembly by binding Mic10 in the IMM-proximal region and Mic60 in the intermembrane space. Taken together, our integrative approach sheds light on the structure and assembly of the MICOS complex.
Rawat, P.; Ramakrishnan, P.; Cardente, N.; Kumar, S.; Greiff, V.; Gromiha, M. M.
Show abstract
Protein aggregation is central to amyloid-related disorders and remains a major developability challenge for protein therapeutics. Over the past two decades, significant advances have been made to predict aggregation-prone regions (APRs) and estimate aggregation propensity in proteins and peptides. In contrast, the prediction of aggregation kinetics has received relatively less attention due to the limited availability and heterogeneity of experimental data. Consequently, aggregation propensities from APR prediction algorithms were widely accepted as a means to predict relative changes in the aggregation kinetics of proteins and mutants. Previous studies have demonstrated, using large-scale datasets, that aggregation propensity shows a weak or inconsistent correlation with aggregation kinetics. In the present study, we have integrated complementary state-of-the-art mechanistic and kinetic prediction tools for protein aggregation into a unified, user-friendly web framework entitled "Amylo-Pipe". Amylo-Pipe also implements practical features that are especially useful for protein engineering, such as gatekeeper-residue mutational scanning to support the design of aggregation-resistant variants. By consolidating multiple prediction tasks in a single interface, Amylo-Pipe enables a more comprehensive assessment of aggregation behavior than APR-only workflows. The web server is freely accessible at: https://web.iitm.ac.in/bioinfo2/amylopipe/.
Malhis, N.; Mehdiabadi, M.; Erdos, G.; Gsponer, J.; Kurgan, L.; Tosatto, S. C. E.; Dosztanyi, Z.; Piovesan, D.
Show abstract
Computational predictors of protein-binding sites within intrinsically disordered regions (IDRs) show highly inconsistent performance across high-quality benchmark datasets. To understand the origins of these discrepancies, we systematically compared predictors across three independent test sets: two CAID datasets updated with the latest DisProt annotations and a composite dataset (DBs) assembled from DIBS, FuzDB, IDEAL, and MFIB. Predictors trained predominantly on DisProt data achieved substantially higher AUCs on the CAID sets but performed poorly on the DBs. In contrast, predictors trained on older, low-quality PDB-based datasets showed balanced performance across all sets, with a slight preference for DBs. Predictors with mixed training exposure displayed intermediate behavior. Through controlled experiments using identical CNN architectures and feature analysis, we demonstrate that the dominant factor driving these performance differences is the intrinsic disorder propensity of the binding sites themselves. Binding residues in DisProt-based datasets exhibit markedly higher average disorder propensity scores than those in PDB-derived datasets. This previously unrecognized selection bias -- literature studies preferentially characterizing more disordered binding sites, while PDB-derived annotations capture less disordered ones -- effectively splits IDR-protein binding sites into two distinct categories. Predictors optimized on one category therefore generalize poorly to the other. Binding-site length and sequence conservation play only minor or negligible roles in explaining the observed inconsistencies. These findings highlight a critical limitation in current benchmarking practices and training strategies for IDR-binding site prediction, underscoring the need for more balanced and disorder-aware reference datasets. Finally, the diagnostic techniques introduced here could prove valuable beyond the specific application examined in this study.
Zhu, Y.
Show abstract
Antibodies provide programmable molecular recognition, whereas enzymes enable repeated chemical transformation. Catalytic antibodies seek to combine these properties within a single protein scaffold. However, conventional approaches based on transition state analogue immunisation, library screening or local mutagenesis provide limited control over the atomic arrangement of catalytic residues. They also frequently produce antibodies that bind substrates without supporting efficient chemical turnover. Recent advances in generative protein design have enabled the construction of antibody complementarity determining regions and the scaffolding of functional motifs under structural constraints. A systematic strategy for transferring experimentally supported enzyme active site geometry into antibody variable domains is still lacking. Here, we present a computational framework that treats antibody and enzyme structures as distinct but complementary inputs. Developable Fv or VHH structures provide the immunoglobulin scaffold. Enzyme complexes containing substrates, products or transition state analogues provide catalytic residues, ligand conformations, metals, cofactors and key water networks. The selected catalytic atoms are mapped into antibody complementarity determining regions, while the surrounding loops are reconstructed using antibody compatible representations and constrained all atom diffusion. Sequence design and structural back prediction are followed by filters for antibody folding, catalytic geometry, ligand positioning, conformational stability and developability. The framework avoids direct fusion of intact enzymes and antibodies. Instead, it transfers only the local geometry required for catalysis. This separation of scaffold selection from catalytic motif selection creates a testable route for determining whether natural enzyme chemistry can be embedded within antibody formats. It also provides a practical basis for evaluating substrate binding, chemical conversion, product release and catalytic turnover as separate design objectives. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=91 SRC="FIGDIR/small/740676v1_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@16ef6d6org.highwire.dtl.DTLVardef@f7c23org.highwire.dtl.DTLVardef@9ee50borg.highwire.dtl.DTLVardef@1cf42d4_HPS_FORMAT_FIGEXP M_FIG C_FIG
Scherlo, M.; Wippermann, E.; Fuertges, T.; Kuenne, R.; Yelboga, A.; Ruetten, F.; Boeckmann, M.; Hoeweler, U.; Rudack, T.
Show abstract
Molecular interactions govern cellular function, making them essential to discover biomolecular mechanisms by unravelling structure-function relationships. The rapid growth of AI-based prediction, experimental determination, and molecular dynamics simulations generates structural data at an unprecedented scale. However, structural information is typically represented as Cartesian coordinates, leaving chemical interactions and conformational relationships largely implicit. We introduce a high-throughput framework transforming structural geometry into a standardized, compact contact space. Moving beyond simple distance cutoffs, it provides a chemically and geometrically informed representation of various residue-residue interactions, their temporal changes, and conformations at residue-level resolution. Our contact-space representation enables systematic comparison and classification even for large-scale analysis. Case studies spanning structure comparison or studies of protein-protein, protein-ligand, protein-RNA, and antibody-antigen complexes, demonstrate how contact-space analysis reveals interaction patterns, identifies key mutation sites, and links structural features to experimental observations. With these and further applications, kontakteUR elucidates biomolecular function and assists targeted protein design, with results suited for further processing by artificial intelligence algorithms.
Eicholt, L. A.; Middendorf, L.
Show abstract
Structure and disorder predictors are increasingly used as decision-grade tools in protein engineering and in the analysis of newly emerged proteins, yet how the current state-of-the-art behaves on sequences outside the well-charted evolutionary space remains poorly characterised. We previously reported that AlphaFold2 confidence and the disorder predictor flDPnn produced discordant predictions for naturally evolved de novo Drosophila proteins and for shuffled sequences. Here, we revisit the comparison with AlphaFold3 and the best-performing disorder predictor PUNCH2 on the same sequence sets together with conserved Drosophila proteins and intrinsically disordered proteins. The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with {beta}-strand fraction, opposite to the conserved and disordered baselines. A class-specific, score-defined driver subset jointly captures the unusual high-pLDDT, high-disorder, low-strand combination and contains 24.5% of de novo, 29.4% of random, 5.1% of conserved, and 1.3% of disordered proteins. Removing this subset normalises the correlations. A held-out classifier trained on architectural and compositional features that were not used in the driver definition recovers the subset, with helix and coil fraction, sequence length, entropy and hydropathy as the strongest predictors. The discordance is therefore not a sequence-class artefact but a localised, compositionally identifiable phenotype that current predictors handle in a non-canonical way - a concrete failure mode that protein designers and others working on sequences remote in sequence space should be aware of when relying on predictor outputs.